Papers with MIT License

4 papers
GraphQL Query Generation: A Large Training and Benchmarking Dataset (2024.emnlp-industry)

Copied to clipboard

Challenge: GraphQL is a powerful query language for APIs, but crafting complex GraphqL queries can be challenging.
Approach: a team of researchers has created a large-scale, cross-domain text-to-GraphQL query operation dataset . the dataset includes 10,940 training triples spanning 185 cross-source data stores and 957 test triples over 14 data stores.
Outcome: The proposed dataset includes 10,940 training triples and 957 test triples over 14 data stores.
Instructional Fingerprinting of Large Language Models (2024.naacl-long)

Copied to clipboard

Challenge: Large language models (LLMs) require considerable cost to train from scratch . fingerprinting is essential to protect intellectual property and to ensure downstream users and developers adhere to their license terms.
Approach: They propose a method for fingerprinting Large language models that implants a private key into the model to generate specific text when the key is present.
Outcome: The proposed method prevents publisher overclaim and maintains robustness against fingerprint guessing and parameter-efficient training.
Building Sentiment Lexicons for Mainland Scandinavian Languages Using Machine Translation and Sentence Embeddings (2022.lrec-1)

Copied to clipboard

Challenge: a simple but effective method to build sentiment lexicons for the three Mainland Scandinavian languages is proposed . a number of experiments with Scandinavian language datasets yield state-of-the-art results using a rule-based sentiment analysis algorithm.
Approach: They propose a simple but effective method to build sentiment lexicons for the three Mainland Scandinavian languages.
Outcome: The proposed method is based on the English Sentiwordnet and a thesaurus in one of the target languages.
NuNER: Entity Recognition Encoder Pre-training via LLM-Annotated Data (2024.emnlp-main)

Copied to clipboard

Challenge: Named Entity Recognition (NER) is a core component of natural language processing, present in a variety of applications such as medical coding, financial news analysis, or legal documents parsing.
Approach: They propose to use Large Language Models (LLMs) to create NuNER, a compact language representation model specialized in the Named Entity Recognition task.
Outcome: The proposed model outperforms similar-sized foundation models in the few-shot regime and is based on a human-annotated dataset.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations